Papers with reference-free metrics
Fusion-Eval: Integrating Assistant Evaluators with LLMs (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Recent studies have employed large language models (LLMs) as reference-free metrics for NLG evaluation, enhancing adaptability to new tasks tasks. |
| Approach: | They propose a method that leverages large language models to integrate insights from various assistant evaluators. |
| Outcome: | The proposed approach achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.7444 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods. |
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment (2023.emnlp-main)
Copied to clipboard
| Challenge: | Conventional reference-based metrics have low correlation with human judgments, especially for open-ended generation tasks. |
| Approach: | They propose to use large language models as reference-free NLG evaluators to assess the quality of NLG outputs. |
| Outcome: | The proposed framework outperforms all previous methods in two generation tasks, and has a Spearman correlation of 0.514 with human on summarization task, and a large variance in human judgments. |
CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation (2022.acl-long)
Copied to clipboard
| Challenge: | Existing reference-free metrics have obvious limitations for evaluating controlled text generation models. |
| Approach: | They propose an unsupervised reference-free metric which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks. |
| Outcome: | The proposed metric has higher correlations with human judgments while obtaining better generalization of evaluating generated texts from different models and with different qualities. |
On the Evaluation Metrics for Paraphrase Generation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for paraphrase generation are not designed for the task, but adopted from other evaluation tasks. |
| Approach: | They propose a new evaluation metric for paraphrase generation that uses reference-based and reference-free metrics. |
| Outcome: | The proposed evaluation metric outperforms existing metrics and is more reliable than reference-based metrics. |
Is Reference Necessary in the Evaluation of NLG Systems? When and Where? (2024.naacl-long)
Copied to clipboard
| Challenge: | Despite recent advances in reference-free metrics, it has not been well understood when and where they can be used as an alternative to reference-based metrics. |
| Approach: | They propose to use reference-free metrics to evaluate NLG systems . they find they have a higher correlation with human judgment and greater sensitivity to deficiencies in language quality . |
| Outcome: | The proposed metrics exhibit higher correlation with human judgment and greater sensitivity to deficiencies in language quality. |
A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing syntactically-controlled paraphrase generation models perform well with human-annotated or well-chosen syntaktic templates. |
| Approach: | They propose a quality-based Syntactic Template Retriever to retrieve templates based on the quality of the to-be-generated paraphrases. |
| Outcome: | The proposed algorithm can generate high-quality paraphrases without sacrificing quality. |
AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset (2025.findings-acl)
Copied to clipboard
| Challenge: | Identifying factors that make ad text attractive is essential for advertising success . identifying the linguistic factors presents a significant challenge because of the intricate interplay between the semantic content and its linguistic expression. |
| Approach: | They propose to use a dataset for ad text paraphrasing that contains human preference data to enable analysis of linguistic factors. |
| Outcome: | The proposed dataset is 20 times larger than v1.0 and contains 16,460 pairs of ad text paraphrase pairs . it shows that human preference and ade- t attractiveness are related . |
Reference-Free Evaluation of Taxonomies (2026.findings-acl)
Copied to clipboard
| Challenge: | Taxonomies are used to classify items, ideas or organisms based on shared characteristics. |
| Approach: | They introduce two reference-free metrics for quality evaluation of taxonomies in the absence of labels. |
| Outcome: | The proposed metrics correlate well with F1 against ground truth taxonomies on five taxonomies and improve hierarchical classification when used with label hierarchies. |
When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in topic segmentation have led to a surge in interest in reference-free metrics, designed to score a hypothesised segmentation of a document without the need to refer to any expert annotation. |
| Approach: | They propose a common framework for reference-free topic segmentation metrics and a new method for the embedding space. |
| Outcome: | The proposed framework outperforms existing metrics based on human annotations while allowing for conversational data to outperformed other metrics. |